speech bubble
Crossing Language Borders: A Pipeline for Indonesian Manhwa Translation
Narasimhan, Nithyasri, Singh, Sagarika
In this project, we develop a practical and efficient solution for automating the Manhwa translation from Indonesian to English. Our approach combines computer vision, text recognition, and natural language processing techniques to streamline the traditionally manual process of Manhwa(Korean comics) translation. The pipeline includes fine-tuned YOLOv5xu for speech bubble detection, Tesseract for OCR and fine-tuned MarianMT for machine translation. By automating these steps, we aim to make Manhwa more accessible to a global audience while saving time and effort compared to manual translation methods. While most Manhwa translation efforts focus on Japanese-to-English, we focus on Indonesian-to-English translation to address the challenges of working with low-resource languages. Our model shows good results at each step and was able to translate from Indonesian to English efficiently.
Context-Informed Machine Translation of Manga using Multimodal Large Language Models
Lippmann, Philip, Skublicki, Konrad, Tanner, Joshua, Ishiwatari, Shonosuke, Yang, Jie
Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding and underdeveloped field and presents complexities even greater than those found in standard translation due to the need to effectively incorporate visual elements into the translation process to resolve ambiguities. In this work, we investigate to what extent multimodal large language models (LLMs) can provide effective manga translation, thereby assisting manga authors and publishers in reaching wider audiences. Specifically, we propose a methodology that leverages the vision component of multimodal LLMs to improve translation quality and evaluate the impact of translation unit size, context length, and propose a token efficient approach for manga translation. Moreover, we introduce a new evaluation dataset -- the first parallel Japanese-Polish manga translation dataset -- as part of a benchmark to be used in future research. Finally, we contribute an open-source software suite, enabling others to benchmark LLMs for manga translation. Our findings demonstrate that our proposed methods achieve state-of-the-art results for Japanese-English translation and set a new standard for Japanese-Polish.
A Comprehensive Gold Standard and Benchmark for Comics Text Detection and Recognition
Soykan, Gürkan, Yuret, Deniz, Sezgin, Tevfik Metin
This study focuses on improving the optical character recognition (OCR) data for panels in the COMICS dataset, the largest dataset containing text and images from comic books. To do this, we developed a pipeline for OCR processing and labeling of comic books and created the first text detection and recognition datasets for western comics, called "COMICS Text+: Detection" and "COMICS Text+: Recognition". We evaluated the performance of state-of-the-art text detection and recognition models on these datasets and found significant improvement in word accuracy and normalized edit distance compared to the text in COMICS. We also created a new dataset called "COMICS Text+", which contains the extracted text from the textboxes in the COMICS dataset. Using the improved text data of COMICS Text+ in the comics processing model from resulted in state-of-the-art performance on cloze-style tasks without changing the model architecture. The COMICS Text+ dataset can be a valuable resource for researchers working on tasks including text detection, recognition, and high-level processing of comics, such as narrative understanding, character relations, and story generation. All the data and inference instructions can be accessed in https://github.com/gsoykan/comics_text_plus.
AI extracts speech bubbles from comic strips
Case in point: Researchers at Google parent company Alphabet's DeepMind recently revealed in an academic paper that they'd developed a system capable of segmenting CT scans with "near-human performance." Now, scientists at the University of Potsdam in Germany have developed an AI segmentation tool for a slightly more cartoony medium: comics. In a paper published on the preprint server Arxiv.org During tests involving a dataset containing speech bubbles with "wiggly tails" and "curved corners," it achieved an F1 score (a measure of a test's accuracy) of 0.94, which the researchers claim is state-of-the-art. "Speech balloons usually consist of a carrier, [a symbolic device used to hold the text,] and a tail connecting the carrier to its root character from which the text emerges. Both tails and carriers come in a variety of shapes, outlines, and degrees of wiggliness," the researchers explain.
Samsung's C-Lab adds character to AI at SXSW
The first concept on show was Toonsquare, which uses AI to convert sentences into cartoons. Like Samsung's AR Emoji, the process starts with a selfie but instead of creating a creepy 3D version of you, it generates a cutesy chibi. A few of us tried this out, and each time the character was a convincing (if unflattering) representation. Once you have your character, you type words into speech bubbles, and the AI will work to discern the emotions in each sentence. It'll then customize the pose and expression of the character, and the formatting of the speech bubble, to match the words.
Like parents from the 1950s, AI still can't understand comics. Here's why
Image recognition has progressed in leaps and bounds over the years. Not too long ago, a challenging recognition task involved asking an AI "Is there a human in this image?" More recently, however, the bar has been raised -- and a new research project carried out at the University of Maryland and University of Colorado has another recognition task in its sights: whether or not an AI can read comic books. In some ways, this is deeply ironic. For a long time, comics were dismissed as a junk medium for kids and barely-literate adults.
Like parents from the 1950s, AI still can't understand comics. Here's why
Image recognition has progressed in leaps and bound over the years. Not too long ago, a challenging recognition task involved asking an AI "Is there a human in this image?" More recently, however, the bar has been raised -- and a new research project carried out at the University of Maryland and University of Colorado has another recognition task in its sights: whether or not an AI can read comic books. In some ways, this is deeply ironic. For a long time, comics were dismissed as a junk medium for kids and barely-literate adults.
Google Will Use Machine Learning To Bring Comic Books To Life - ARC
As an avid comic book reader, I know that the outside world falls into two categories … people who love comics and those who think that comic book fans don't know how to read a real book. I have had numerous discussions over the years with friends, family and significant others who tell me that comics are for kids or that there is no literary value in graphic novels. My standard response is to hand them a copy of Alan Moore's Watchmen and tell them to read it. If they still don't think that comic books have value, then I quietly point out that most blockbuster movies since 2008 have been based on characters from either the Marvel or DC universes (well, maybe not DC so much). And if they are remain unconvinced, then I make a mental note to not mention comics again.
Google now uses machine learning to make reading comics on phones easier
Plenty of people still like to read their comics on paper, but increasingly, phones and tablets are the devices of choice for keeping up with the Justice League. Last year, Google introduced a new reading experience for comics in its Play Books store for Android that makes it easier to follow along with the story. Today, the company is launching yet another update to the comics reading experience -- this time with a focus on making the speech bubbles in comics more readable on small devices. As Google's Head of Product for Play Books Greg Hartrell told me, the team looked at the feedback it got from the last update. While readers liked the new reading experience, they complained that it was still too hard to read the text on a small screen.